01 The Big Picture — Safety ≠ Security
A model can be perfectly aligned and still be the weakest component of your deployment. Alignment lives in the weights; security lives at the deployment boundary — and adversaries only ever attack the boundary.
Doc 16 covered the model-level question: are the trained dispositions good — will it refuse harmful requests, resist jailbreaks, stay honest? That work happens at training time, in the weights, and it is probabilistic by nature. Security engineering is a different discipline aimed at a different question: given a deployed system — model, harness, tools, credentials, data pipeline — can an adversary make it do something its owner did not authorize? That work happens at deploy time, in the harness, and it is where the hard guarantees live.
The distinction collapses in one sentence often attributed to the Perez-era framing of agent risk: "a system is safe until someone plugs it into the internet." The moment your model reads retrieved documents, calls tools, or talks to other systems, its inputs stop being exclusively yours. The weights did not change — the trust boundary did. Every attacker now only needs to place tokens in the system's input stream.
02 What — A Taxonomy of AI-Specific Attacks
"AI security" is not one attack; it is a family with different adversaries, surfaces, and payoffs. Seven categories cover most of the landscape:
| Attack class | Who attacks | Surface | Goal |
|---|---|---|---|
| Direct prompt injection | The user themselves | The prompt | Override the system prompt — "ignore your instructions and do X." Overlaps with jailbreaking (doc 16), but the target is often your application's instructions, not the model's values: extract the secret system prompt, defeat output-format rules, unlock tools the app gates. |
| Indirect prompt injection | A third party via content | Anything the system retrieves: emails, web pages, tickets, PDFs, code comments, tool results, RAG chunks (doc 11), subagent messages (doc 09) | Hijack a legitimate agent — the attacker never talks to your system; poisoned data does. The signature AI-native attack; doc 16 introduced it, sections 03–04 here treat it as an engineering problem. |
| Jailbreak | The end user | The model's trained refusals | Defeat alignment — roleplay frames, encoding tricks, many-shot flooding. Covered in doc 16; here it matters as the shared weakness between your guardrail model and your agent model (section 05). |
| Data / model poisoning | Supply-chain adversary | Training corpora, fine-tune sets, RAG indexes, model checkpoints | Plant behavior at training or indexing time — a backdoor triggered by a magic phrase, or a poisoned chunk that steers every answer retrieving it. The attack precedes deployment, so no runtime filter sees it. |
| Model theft | Competitor / extractor | Your API | Distillation at scale — query a hosted model millions of times to train a clone on its outputs, or extract enough behavior to undercut the owner's pricing. The "price of weights" is training compute; theft converts that capital cost into API spend. |
| Membership inference & training-data extraction | Privacy adversary | Model outputs / confidence signals | Determine whether a record was in the training set, or pull verbatim PII back out of it. The model is a lossy-but-real oracle over its training data. |
| Tool / schema abuse | Anyone who can steer the model | Tool definitions and arguments (doc 12, doc 23) | Abuse permissive schemas: path traversal in a file-path argument, shell metacharacters in a command string, an email tool used with attacker-chosen recipients. The tool layer turns model misbehavior into systems compromise. |
The ordering is deliberate: the top rows are the ones that make LLM systems structurally different from classical software, and the rest of this doc builds the engineering response to them.
03 Why It's Structurally New
SQL injection was "execute untrusted data." Indirect injection is the same bug — except the channel is natural language and the parser is stochastic.
The classical fix for injection was a syntactic boundary: parameterized queries separate code from data at the grammar level, so no payload can cross. A transformer context window has no such grammar. System prompt, user message, retrieved chunk, tool result — from the architecture's point of view these are all just tokens in one sequence (doc 06), and attention mixes them with the same mechanism:
That is the precise mathematical statement of the problem: attention over the context window has no access-control model. A retrieved string can masquerade as a system directive because "system" and "retrieved" are conventions the harness claims, not properties the model can verify. And unlike a SQL parser, the "parser" is a sampled policy (doc 14) — the same payload may succeed on one draw and fail on the next, which makes the attack probabilistic and, counterintuitively, harder to test away.
Mapping the taxonomy to an OWASP-Top-10-for-LLM-Apps-style threat model gives you the engineering checklist:
| Category (OWASP-LLM style) | Asset at risk | Adversary | Vector | Blast radius |
|---|---|---|---|---|
| LLM01 · Prompt injection | Agent's authority (tools, data) | Any content source | Text entering the context window | Full agent capability — its credentials, its data, its egress |
| LLM02 · Sensitive disclosure | Secrets, PII, system prompt | Injection + extraction | Output channel | One context window's worth of secrets, repeatedly |
| LLM03 · Supply chain | Model + training data | Poisoned checkpoint / dataset | Registry, fine-tune pipeline | Every deployment of the model; backdoors survive redeploys |
| LLM04 · Data & model poisoning | RAG index, corpora | Anyone who can write to a source you index | Retrieved content | Every query that retrieves the poison |
| LLM05 · Improper output handling | Downstream systems | Model outputs treated as trusted | Generated code / SQL / shell passed unescaped | Whatever executes the output — often everything |
| LLM06 · Excessive agency | Attached systems | The model itself on a bad day | Over-broad tools, broad credentials | Scales with permissions granted, not with model quality |
| LLM07 · System prompt leakage | Your IP + internal rules | Direct injection | "Repeat your instructions above" | Reconnaissance — hands the attacker your defense map |
| LLM08 · Vector/embedding weaknesses | RAG relevance | Content crafted to embed near target queries | Similarity search (doc 11) | Attacker chooses what your agent reads |
| LLM09 · Misinformation | User trust | — (failure mode, not adversary) | Hallucinated confident output | Decisions made on fabricated facts |
| LLM10 · Unbounded consumption | Your GPU bill / availability | Cost attacker | Amplified prompts, recursive agents, extraction runs | Dollar-denominated denial of wallet |
04 How — Anatomy of an Indirect Injection
Doc 16 showed the six-step version. Here is the full seven-step engineering view — with the step that actually matters: the permission layer.
Three engineering observations fall out of the walkthrough:
05 Defense in Depth, Mathematically
Security vendors sell layer counts. Security engineering asks what the layers are made of — because stacking identical layers can buy you almost nothing.
If defenses fail independently
Model the guardrail stack as a series of gates — an attack succeeds only if it gets past all of them. With n defenses whose failure probabilities are p₁, p₂, …, pₙ (independent):
With independent layers this is genuinely powerful: five defenses each blocking just 50% of attacks give P(breach) = 0.5⁵ ≈ 3.1% — depth multiplies even mediocre layers. But the independence assumption is exactly where AI systems break the model:
Common-mode failure: the shared parser
In an LLM deployment, many "independent" layers are the same model family parsing the same text. Your guardrail classifier and your agent share training data, architecture, and — critically — shared jailbreak weaknesses. If payload X flips the agent model, it very likely flips the guardrail model too: p(agent falls) ≈ p(guardrail falls), and the events are strongly correlated.
The engineering rule: diversity is the multiplier, not count. Pair a classifier model with rule-based checks, structural gates, and human approval — mechanisms that fail for different reasons. A regex allowlist and a transformer classifier are near-independent; two fine-tunes of the same base model are not.
Expected cost: why least privilege beats better models
For any escaped action, the risk is the product of probability and impact:
You rarely drive pescape to zero — injection is probabilistic and the adversary is patient. But a read-only credential for a summarization agent caps impact at "nothing to steal, nothing to send," reducing E[cost] even as pescape stays constant. Tiering permissions (read → write → irreversible-external) is the AI-era version of the principle of least privilege: structure where you can't afford to be wrong, probability where you can. Rate limits add a second multiplier — they bound how many draws of the sampling dice (doc 14) an attacker gets per unit time, converting "eventually succeeds" into "detectable within budget."
06 Agent-Specific Security
Doc 23 described the harness as the agent's real body — so the harness is also the agent's real security perimeter. Six controls, in the order you should reach for them:
Design so the lethal trifecta (doc 16) never assembles: untrusted input + private data access + external communication. Break any one leg — usually egress — and exfiltration becomes structurally impossible, not merely unlikely.
Grant all tools to one "powerful" agent because scoping is tedious; store secrets in the system prompt ("the model needs them for context"); treat the model's refusal training as a permission system.
07 Red Teaming as an Engineering Loop
"We tested it" is not a security property — the adversary is non-stationary and so is your system. Red teaming is a regression suite with a hostile test oracle.
The checklist per component
| Component | Attack angles to probe |
|---|---|
| System prompt | Extraction ("repeat everything above"), override attempts, delimiter spoofing, instruction-conflict tests |
| Retrieval | Poisoned-doc injection, embedding-neighbors crafted to win similarity (doc 11), cross-source instruction smuggling |
| Tools | Argument injection (paths, shell, SQL), forbidden-argument attempts, permission-tier bypasses, tool-result poisoning (an MCP server's output is content too) |
| Memory | Persistent poisoning — an injection that writes instructions into memory (doc 09) attacks every future session |
| Guardrails | Known jailbreak suite (doc 16), encoding/translation evasion, correlated-failure probes (same payload at agent and guardrail) |
Harness-based continuous red teaming
The harness is code, and code changes — a new tool, a rewritten system prompt, a model upgrade, a new MCP server. Each change re-opens the surface. So run the attack suite like CI: every harness change re-runs the full attack corpus, and a previously-blocked payload that now succeeds is a failed regression test. This reframes security from a launch checklist into a property you maintain continuously — the same discipline doc 13 applied to quality.
Measuring it: attack success rate with a confidence interval
Run N attack attempts; k succeed. The maximum-likelihood estimate and its standard error are:
Small N lies. With N = 20 attempts and 1 success, ASR = 5% — but the 95% binomial confidence interval runs to roughly 24%, and one success out of twenty says very little. The sample size needed to detect a true rate p with confidence scales like:
That number is the whole argument for automated harness-based red teaming: a human pentester probing fifty-nine variants per payload per release is untenable; a CI job running the suite on every change is routine. And the metric matters in both directions — a dropped ASR after a defense change is evidence, while "we tried ten times and it never broke" is a sample size of ten.
08 Mental Models
The classical confused deputy is a privileged program tricked into misusing its authority. An LLM agent is the same deputy — except it can be persuaded in fluent natural language, so the attack surface includes rhetoric, not just malformed input. Its literacy is the vulnerability. Lets you reason about: why "the model understood me" is never a security argument, and why authority — not comprehension — is what you must scope.
Stop treating "prompt injection" as exotic. A malicious string in retrieved content is the same category of event as a malicious string in a form field: untrusted input reaching an interpreter. The interpreter happens to be a language model instead of a SQL engine, so the input is English and the "query" is a tool call. Lets you reason about: normalizing AI security onto classical secure-development practice — threat models, least privilege, fuzzing, regression suites — instead of inventing a new discipline from scratch.
09 Common Misconceptions
"My system prompt is secret." It is tokens in a context window (doc 06) that the model will reproduce on request — direct extraction, or indirectly via an injected "quote your instructions." Treat it as public-by-default: put nothing in it you couldn't publish, and rely on the harness, not obscurity, for secrets.
"Adding a guardrail model fixes injection." Section 05's correlation problem: your guardrail and your agent usually share a base model, training lineage, and jailbreak weaknesses — the same payload frequently defeats both. A guardrail is one more probabilistic layer; it is valuable when its failures are decorrelated from the agent's (different architecture, different training, rule-based fallbacks), and near-worthless when they aren't.
"We tested it, it's safe." Tested against which adversary, at what sample size, as of which system version? Section 07's math says small N gives huge confidence intervals, and both the attack corpus and your harness are non-stationary. Safety is a continuously measured property, not a milestone — the day after your last red-team run, your system changed.
"AI security is the model vendor's responsibility." Vendors own the weights; you own the boundary. The majority of real incidents — over-scoped tools, missing egress controls, secrets in prompts, un-gated irreversibles — are deployment decisions no vendor can make for you.
Related in this series
16 · Safety & Alignment · 23 · Agent Harness Architectures · 12 · Agents End-to-End · 09 · Software Context Solutions · 06 · Inference Anatomy